Papers with context-aware evaluation
On Measuring Context Utilization in Document-Level MT Systems (2024.findings-eacl)
Copied to clipboard
| Challenge: | Current studies on document-level translation evaluation focus on sentence-level models which are inadequate for capturing improvements in discourse phenomena. |
| Approach: | They propose to complement accuracy-based evaluation with measures of context utilization. |
| Outcome: | The proposed model can be used to handle context-dependent discourse phenomena using an automatic annotation tool. |
Chimera: Compositional Jailbreak Attacks on LLMs via Judgment-Driven Search over Heterogeneous Strategies (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for evaluating large language models face two limitations: they explore homogeneous transformations in isolation and rely on brittle judgment metrics that misclassify non-refusal hallucinations as successful attacks. |
| Approach: | They propose a framework that generates compositional jailbreak attacks via judgment-driven search over heterogeneous strategies. |
| Outcome: | The proposed framework generates compositional jailbreak attacks over heterogeneous strategies . strongREJECT++ improves attack success rates and transferability compared to state-of-the-art . |